Papers with English-centric corpora

2 papers
Accelerating Multilingual Language Model for Excessively Tokenized Languages (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown a significant degree of multilingual proficiency on a variety of tasks in multiple languages.
Approach: They propose a framework to fine-tune a language model head and fine-track it while preserving its performance.
Outcome: The proposed framework increases the generation speed by 1.7 while maintaining the performance of pre-trained multilingual models on target monolingual tasks.
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit strong multilingual performance despite training on English-centric corpora.
Approach: They propose to use Romanization as a potential bridge in multilingual processing . they propose to encode semantic concepts similarly across native and Romanized scripts .
Outcome: The proposed model encodes semantic concepts across native and Romanized scripts, suggesting a shared underlying representation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations